Papers with sign language translation
Sign Language Translation with Sentence Embedding Supervision (2024.acl-short)
Copied to clipboard
| Challenge: | State-of-the-art sign language translation systems facilitate learning through gloss annotations when available at scale. |
| Approach: | They propose to use sentence embeddings of the target sentences at training time that take the role of glosses to supervise the learning process. |
| Outcome: | The proposed method significantly outperforms gloss-free approaches on German and American sign languages and with mono- and multilingual sentence embeddings and translation systems. |
Considerations for meaningful sign language machine translation based on glosses (2023.acl-short)
Copied to clipboard
| Challenge: | In machine translation, sign language translation based on glosses is becoming more popular . limitations of glossed approaches are not discussed in a transparent manner, and there is no common standard for evaluation. |
| Approach: | They propose to use a gloss-based approach to evaluate machine translation results . they propose to include realistic datasets, stronger baselines and convincing evaluation . |
| Outcome: | The proposed approach is based on a neural gloss translation model. |
Machine Translation between Spoken Languages and Signed Languages Represented in SignWriting (2023.findings-eacl)
Copied to clipboard
| Challenge: | Yin et al. ( 2021) calls for including sign language processing (SLP) in natural language processing research. |
| Approach: | They propose to use a sign language writing system to parse, factorize, decode and evaluate signed languages. |
| Outcome: | The proposed method achieves over 30 BLEU in a bilingual setup and over 20 BLUE in two multilingual setups. |
Signer Diversity-driven Data Augmentation for Signer-Independent Sign Language Translation (2024.findings-naacl)
Copied to clipboard
| Challenge: | Existing methods for sign language translation (SLT) rely on signer identity labels, which is often impractical and costly in real-world applications. |
| Approach: | They propose a signer diversity-driven data augmentation method that can generalize to signers not encountered during training. |
| Outcome: | The proposed method achieves state-of-the-art results without relying on signer identity labels. |
Sign Language Production With Avatar Layering: A Critical Use Case over Rare Words (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing vision-based sign language production approaches suffer from out-of-vocabulary (OOV) and test-time generalization problems. |
| Approach: | They propose an avatar-based sign language production system that generates sign language videos from spoken language expressions. |
| Outcome: | The proposed system achieves higher BLEU-4 and higher ROUGE-L scores on a new Korean-Korean sign language dataset. |
GRAMMAR-LLM: Grammar-Constrained Natural Language Generation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to fine-tuning and prompting are insufficient to ensure compliance with predefined taxonomies, syntactic structures, or domain-specific rules. |
| Approach: | They propose a framework that integrates formal grammatical constraints into the decoding process to enforce syntactic correctness in linear time while maintaining expressiveness in grammar rule definition. |
| Outcome: | The proposed framework enforces syntactic correctness in linear time while maintaining expressiveness in grammar rule definition. |
Explore More Guidance: A Task-aware Instruction Network for Sign Language Translation Enhanced with Data Augmentation (2022.findings-naacl)
Copied to clipboard
| Challenge: | Existing studies focus on the recognition step, while paying less attention to sign language translation. |
| Approach: | They propose a task-aware instruction network, namely TIN-SLT, for sign language translation, by introducing the isntruction module and the learning-based feature fuse strategy into a Transformer network. |
| Outcome: | The proposed system outperforms existing solutions on two benchmark datasets, PHOENIX-2014-T and ASLG-PC12, and outperformed previous best solutions by 1.65 and 1.42 in terms of BLEU-4. |
Improvement in Sign Language Translation Using Text CTC Alignment (2025.coling-main)
Copied to clipboard
| Challenge: | Current sign language translation (SLT) approaches rely on gloss-based supervision with Connectionist Temporal Classification (CTC) limiting their ability to handle non-monotonic alignments between sign language video and spoken text. |
| Approach: | They propose a method that integrates CTC/Attention with the attention mechanism during decoding and integrates it with the sign language video and spoken text. |
| Outcome: | The proposed method outperforms the pure-attention baseline and achieves comparable results to state-of-the-art methods. |
Prior Knowledge and Memory Enriched Transformer for Sign Language Translation (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for sign language translation do not explore all of them . visual and textual understanding and additional prior knowledge learning are challenging . |
| Approach: | They propose a method which integrates auxiliary information into vanilla transformer for SLT . they use visual-textual context information and additional auxiliary knowledge of a word . |
| Outcome: | The proposed method improves the understanding of sign language videos with visual and textual understanding and additional prior knowledge learning. |
Dynamic Feature Fusion for Sign Language Translation Using HyperNetworks (2025.findings-naacl)
Copied to clipboard
| Challenge: | Using RGB and keypoint streams, sign language translation is highly dependent on the brain's ability to process color, shape, and motion simultaneously. |
| Approach: | They propose a hypernetwork-based fusion method that extracts salient features from RGB and keypoint streams and introduces self-distillation and SST contrastive learning to maintain feature advantages while aligning the global semantic space. |
| Outcome: | The proposed method achieves state-of-the-art performance on two public sign language datasets, reducing model parameters by about two-thirds. |
Reconsidering Sentence-Level Sign Language Translation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Historically, sign language machine translation is framed as a sentence-level task . however, there are known intersentential dependencies that are impossible to resolve in isolation. |
| Approach: | They propose a human baseline for sign language translation that substitutes a person into the machine learning task framing instead of providing the entire document as context. |
| Outcome: | The proposed human baseline for sign language translation shows that deaf signers can only understand key parts of the clip in light of additional discourse-level context. |
Open-Domain Sign Language Translation Learned from Online Video (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing work on sign language translation has focused mainly on data collected in controlled environments or domains, which limits its applicability to real-world settings. |
| Approach: | They propose to use sign search as a pretext task and fusion of mouthing and handshape features to improve sign language translation in real-world settings. |
| Outcome: | The proposed techniques produce consistent and large improvements over baseline models based on prior work. |
Towards Privacy-Aware Sign Language Translation at Scale (2024.acl-long)
Copied to clipboard
| Challenge: | Existing sign language training systems require detailed and time aligned annotations to be effective. |
| Approach: | They propose a two-stage framework for privacy-aware SLT at scale that leverages self-supervised video pretraining on anonymized and unannotated videos followed by supervised SLT finetuning on a curated parallel dataset. |
| Outcome: | The proposed framework outperforms baselines on the How2Sign dataset and achieves state-of-the-art finetuned and zero-shot gloss-free SLT performance. |
Continual Learning in Multilingual Sign Language Translation (2025.naacl-long)
Copied to clipboard
| Challenge: | Despite the low translation quality of sign language, many machine learning approaches are still in its infancy. |
| Approach: | They propose to use continual learning for mul- tilingual SLT to improve translation quality. |
| Outcome: | The proposed methods outperform baseline and fine-tuning approaches in sign language translation. |
JWSign: A Highly Multilingual Corpus of Bible Translations for more Diversity in Sign Language Processing (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing sign language datasets are limited and skewed towards high-income sign languages, mainly those from high-risk countries. |
| Approach: | They propose a large and highly multilingual dataset for sign language translation: JWSign. |
| Outcome: | The proposed dataset consists of 2,530 hours of Bible translations in 98 sign languages, featuring more than 1,500 individual signers. |
Gloss-Free End-to-End Sign Language Translation (2023.acl-long)
Copied to clipboard
| Challenge: | a study of sign language translation without gloss annotations focuses on the problem of gloss annotation . gloss annotation is hard to acquire, especially in large quantities, and limits the domain coverage of translation datasets . |
| Approach: | They propose a gloss-free end-to-end sign language translation framework to solve this problem . gloss annotations are hard to acquire, especially in large quantities, they argue . |
| Outcome: | The proposed framework improves sign language translation performance on large-scale datasets . gloss annotations are hard to acquire, especially in large quantities . |
Unsupervised Sign Language Translation and Generation (2024.findings-acl)
Copied to clipboard
Zhengsheng Guo, Zhiwei He, Wenxiang Jiao, Xing Wang, Rui Wang, Kehai Chen, Zhaopeng Tu, Yong Xu, Min Zhang
| Challenge: | Experimental results on the BBC-Oxford Sign Language dataset reveal that USLNet achieves competitive results compared to supervised baseline models. |
| Approach: | They propose an unsupervised sign language translation and generation network that learns from abundant single-modality data without parallel sign language data. |
| Outcome: | The proposed model achieves competitive results compared to baseline models on the BBC-Oxford Sign Language dataset and Open-Domain American Sign Language data. |
Korean Disaster Safety Information Sign Language Translation Benchmark Dataset (2024.lrec-main)
Copied to clipboard
Wooyoung Kim, TaeYong Kim, Byeongjin Kim, Myeong Jin MJ Lee, Gitaek Lee, Kirok Kim, Jisoo Cha, Wooju Kim
| Challenge: | Sign language is a crucial means of communication for deaf communities. |
| Approach: | They propose to refine Korean sign language translation datasets and release them . they show baseline performance varies depending on tokenization method applied to gloss sequences . |
| Outcome: | The proposed dataset outperforms baseline and spoken language tokenization methods. |
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing work on sign language video processing focuses on the face, hands and body posture of the signer. |
| Approach: | They propose to learn the handshapes and rich facial expressions of sign languages in a self-supervised fashion by learning from individual frames rather than video sequences. |
| Outcome: | The proposed model is more efficient than previous work on sign language pre-training. |
SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for sign language processing have relied on task-specific models, limiting the potential for transfer learning across tasks. |
| Approach: | They propose a self-supervised contextual representation model that adapts masked token prediction objectives to multi-stream visual sign language input. |
| Outcome: | The proposed model adapts masked token prediction objectives to multi-stream visual sign language input, learning to predict multiple targets corresponding to clustered hand, face, and body pose streams. |
Can Small Vision–Language Models Perform Sign Language Translation? (2026.findings-acl)
Copied to clipboard
| Challenge: | Vision-Language Models (VLMs) have shown strong generalization across multimodal tasks, but their capacity to handle sign language translation (SLT) remains unclear. |
| Approach: | They propose entity- and semantics-aware metrics tailored for SLT to evaluate their performance. |
| Outcome: | The proposed metrics highlight the limitations of general-purpose VLMs to SLT, unlike their applicability in other tasks. |
J-Shuwa: A Large-Scale Web-Collected Japanese Sign Language-Japanese Parallel Corpus (2026.findings-acl)
Copied to clipboard
| Challenge: | Japanese Sign Language (JSL) is a low-resource sign language that has received limited attention in the AI community due to the lack of large-scale, publicly available parallel corpora. |
| Approach: | They propose a large-scale JSL-Japanese parallel corpus constructed from YouTube videos with hard-coded subtitles and closed captions. |
| Outcome: | The proposed model is effective for training models and can be used for future research across a wide range of tasks. |